Papers with debiased reward model
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF (2026.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to tackling length bias are limited by their complexity or lack of a linear length-reward relation. |
| Approach: | They propose a framework that learns and corrects underlying bias patterns by fitting a length-reward relationship into a reward model. |
| Outcome: | The proposed framework improves length-controlled win rate and reduces verbosity without compromising performance. |